[Refactor][Quantization]KVCache quantization dtype specified by the --kv-quant-dtype parameter - #15514
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request refactors the KV cache and indexer dtype handling for Ascend-based models. It introduces robust mechanisms to ensure correct dtype resolution (specifically for 'fp8' and 'int8') across different vLLM versions, including a platform-level patch to support 'int8' indexer caches for models like DeepSeek V4. These changes improve kernel compatibility and ensure consistent configuration validation. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Support int8 indexer KV cache dtype and pluggable fp8 kv-cache-dtype fallbackSuggested PR Summary:
### What this PR does / why we need it?
This PR introduces support for `int8` indexer KV cache dtypes, which are required by DeepSeek V4 and similar Ascend sparse-attention models. It implements a platform patch to dynamically widen the Pydantic validation schema for `AttentionConfig.indexer_kv_dtype` to accept `"int8"`. It also adds a worker patch to map `"fp8"` to `torch.float8_e4m3fn` on older vLLM releases lacking the pluggable `register_kv_cache_dtype` mechanism.
However, several critical issues were identified in the review:
- **Crash in non-quantized models**: `enable_fa_quant` raises an empty `ValueError("")` instead of returning `False`, which will crash model initialization for non-quantized models.
- **AttributeErrors**: `self.vllm_config` is accessed in `AscendConfig` and `DeepseekV4IndexerCache` where it is not set or should be accessed via the passed parameter.
- **Uninitialized attributes**: In `sfa_v1.py` and `model_runner_v1.py`, cache dtype attributes are left uninitialized if `indexer_kv_dtype` is `"bf16"` while `enable_sparse_sfa_c8` is `True`.
- **Unconditional call**: `kv_cache_dtype_str_to_dtype` is called unconditionally in `mla_v1.py` which can cause runtime errors for non-quantized models.
### Does this PR introduce _any_ user-facing change?
Yes, users can now specify `--attention_config.indexer_kv_dtype int8` for DeepSeek V4 models on Ascend.
### How was this patch tested?
No tests were provided in the diff. It is recommended to add integration tests for DeepSeek V4 with `int8` indexer KV cache.|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
7b1f42c to
247ddc2
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
1 similar comment
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
7954b8a to
19e7904
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
f6bdb8a to
a253c4f
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
1 similar comment
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
|
/nightly Qwen3-32B-W8A8C8-A3 DeepSeek_V31_W4A4C8_A5 DeepSeek_V4_Flash_A5 GLM5_1_W4A4_A5
|
|
/nightly GLM5_1_W4A4_A5
|
### What this PR does / why we need it? Upgrade the supported vLLM release dependency from **v0.28.0 to the official v0.29.0**, alongside the fixed vLLM main commit inherited from #16216 (already merged). | Reference | Commit | |---|---| | vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` | | PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` | | Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` | | Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a` | | Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` | - Use common implementations where v0.29.0 and the fixed main share KV-cache layouts, Mamba copy/group APIs, PCP handling and speculative-decoding contracts. - Retain explicit `vllm_version_is("0.29.0")` branches for contracts that still differ, including RoPE, scheduler block snapshots, InputBatch, ReplaySSM, KV zeroing and DSpark PP handling. - Remove obsolete v0.28.0 compatibility and adapt existing test fixtures. Version detection uses package versions and the explicit `VLLM_VERSION` override, with local-version suffix handling and UT environment isolation retained; no hard-coded release-SHA inference remains. - Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving other validations and EPLB platform binding. Rebuild dependent Pydantic schemas so nested configuration validation uses the patched validator. The global patch documentation records its rationale and removal criteria. - Preserve #15747's Spec+PP protocol/partition handling after rebase; use the common exact-release selector and the real function-local DSpark sharing import. This release-only upgrade does not require a new main old-to-new interface scan. The latest rebase incorporates [#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files), which reverted #16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The now-unnecessary V4.1 drafter import gate and tuple annotation have been removed. Other release adaptations, including #15747 Spec+PP handling, remain. Detailed contract evidence is retained below for review. <details> <summary>Per-file adaptations and exact upstream evidence</summary> #### Per-file adaptation ledger Evidence IDs refer to the exact source contract and upstream diff table below. Every row is syntax checked; branch-normalized AST comparison confirms unchanged main function bodies except the version identity helper and the explicitly retired propose argument/type annotations. Both supported versions completed CPU and NPU execution as recorded below. | Ascend file / symbols | Disposition and upstream evidence | Verification | |---|---|---| | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` / `_load_dspark_model_with_target_quant`; `tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased [Ascend #15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a) (`82df9d871`), including manual PP partition masking. vLLM [#52809 diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share` inside the loader on both supported pins, so retain the earlier release fix: patch `eagle_utils`, never read/patch a nonexistent `dspark_utils._should_share`. v0.29 keeps its global PP guard; main has native PP via [#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks the real function-local import, absent module alias, and restoration on success/failure for both lanes; partition and PP assertions retained. Ruff/syntax pass; actual CPU/NPU pending. | | `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`; `tests/ut/worker/v2/test_pp_utils.py` | #15747 added broad 0.28/0.29 routing and an obsolete 0.28 dev-build recognition path. For the two supported pins, #50514 exists only on fixed main. Route through `vllm_version_is("0.29.0")`; retain package/local-suffix and explicit environment-override semantics. No release-SHA inference or third release lane. | Existing UTs exercise the real uncached version helper with monkeypatch isolation, local suffix, fixed-main dev string, explicit override, and non-target versions. Isolated routing checks and Ruff/syntax pass; full CPU UT pending. | | `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`, `_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`, `_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436; use shared standardized layouts and retain the v0.29.0 InputBatch gate. The rebase preserves vllm-ascend #16043's MTP copy tracking while removing only legacy v0.28.0 allocation paths. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` | #52839; both lanes use the common PCP import. Rebase keeps current upstream graph code and removes only the v0.28.0 import branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/recompute_scheduler.py`<br>`module imports/dispatch`, `schedule` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`, `AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker` | #52615; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__` | #53614; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module imports/dispatch`, `_get_max_layers_per_page_size`, `_ascend_max_memory_usage_bytes_from_groups`, `_ascend_get_kv_cache_config_from_groups` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` | v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it after KV binding. Evidence: [#52506, `adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c). Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. | Failure reproduced on v0.29.0; Historical release/main NPU validation passed; current-head CI pending | | `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module imports/dispatch` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant` | #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports `_should_share` locally from Eagle utilities; fixed main removes the PP guard. Keep the release `get_pp_group` patch, share through the common Eagle utility, and delete the obsolete v0.28.0 `dspark_utils._should_share` patch. | Exact release failure reproduced; Historical release/main CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`, `__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`, `register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`, `_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`, `_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers | #51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and import the NaN helpers directly because v0.29.0 and fixed main expose the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static checked; historical dual-version CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes `pcp_manager` common to both lanes. The rebase preserves vllm-ascend #16409's host-parameter-update revert and removes only the obsolete v0.28.0 capture branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`, `_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/block_table.py`<br>`__init__`, `init_block_table_layout_tensors`, `compute_slot_mappings` | #51718, #51031; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`, `prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`, `execute_model` | #50514, #54436, #52506, #55212, #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch` | #54282; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` | #49811; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module imports/dispatch`, `propose` | #53694, #52188; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`, `wake_up` | #51718, #53508; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | #### Exact upstream evidence | Upstream change | Full commit SHA / direct diff | Actual supported contracts and branch decision | |---|---|---| | #51718 | `8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Both use layers/layer_stride/block_stride/offset, tokens_per_state, CircularBufferSpec and standardized backing; retire shared_by/compress_ratio allocation branches. | | #52839 | `58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8) | Both import PCP operations from vllm.v1.attention.ops.pcp. | | #53896 | `e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec) | Both use Mamba copy-function dictionaries and unwrap UniformTypeKVCacheSpecs. | | #53106 | `1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21) | Both use WeightsMapper instead of AutoWeightsLoader skip_prefixes/skip_substrs. | | #53906 | `98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Only pinned main has the optional MLA storage_block_size dataclass field; release keeps the Ascend derived property, using tokens_per_state. | | #56078 | `719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417) | Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned main uses mrope_num_dims and unified RoPE. | | #52615 | `138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045) | Release uses num_blocks/kv_bytes_per_block; main uses num_chunks/kv_bytes_per_chunk. | | #51358 | `6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release now has boundary_state_offloads and KVConnectorBlockState; remove partial_tail_offloads plumbing. | | #54853 | `0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release constructor takes block_ids snapshots; main takes req_ids and resolve_block_ids. Keep exact release snapshot membership and main lazy-resolution membership. | | #53614 | `144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725) | Only main configures drop_eagle_checkpoint_block for replay-aligned Mamba checkpoints. | | #50514 | `d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Release retains module-level PP/share symbols and Ascend PP workaround; main has the subsequent PP integration. | | #52809 | `91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark module binding to a function-local import from Eagle utilities. The shared Eagle patch remains effective; the old DSpark-module read/write must be removed. | | #54436 | `6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902) | Release InputBatch requires max_seq_len_np; main removed it. | | #52506 | `adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Only main accepts valid_dummy_state_slots/valid_state_slots capture arguments. | | #55212 | `83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Release prepares DCP local sequence lengths before partitioning; main initializes DCP metadata afterwards. | | #53515 | `b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127) | Both accept padded_num_tokens for persistent PCP input buffers. | | #53869 | `b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212) | Both accept pcp_manager during graph capture. | | #51031 | `0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1) | Both distinguish KV and kernel block sizes during DCP slot mapping. | | #54282 | `fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8) | Both gumbel sampling APIs include is_drafting. | | #52188 | `d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0) | Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. | | #53694 | `5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6) | Both propose APIs take DPSyncState; remove the obsolete token-count argument and retain replicated-PCP synchronization. | | #49811 | `01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e) | Both support extract_hidden_states on MRV2; remove old unsupported dispatch/skip. | | #53508 | `479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29) | Both remove post_kv_cache_wake_up; retire release-only call. | | #52494 | `3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1) | Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain release exclusion. | | #52861 | `b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28) | Both include DeepseekV32MTPModel in the two-hidden-state architecture set. | | #54713 | `b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3) | Only main takes replay_boundaries in compressed-prefix hit lookup; preserve release calls without that keyword. | | #42785 | `442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18) | Only main capture callers pass axis_keys; preserve the existing Ascend rejection of nonempty axes. | | #52358 | `8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Both ExecuteModelState have dp_sync; only main has cudagraph_stats. | | #52789 | `9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4) | Release already has mamba_has_prefill_checkpoint_blocks; later main also has fine-grained prefix-cache state. | | #51251 | `7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | v0.29.0 and main expose ec_manager_config; retire the old release-only ScoreEncoder configuration skip. | | #53240 | `b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c) | v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache groups; retire the old release-only replay skip. | | #53853 | `e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both delegate PCP compatibility validation to the PCP manager. | | #53183 | `4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both expose the V1 unsupported-feature helper used by the existing Ascend MRV1 feature filter. | #### Additional inherited contracts | vllm-ascend change | Why it is required | Upstream cause and direct link | Lane | Verification | |---|---|---|---|---| | `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec, list[int]]` from `get_mamba_groups` and both initialize `recoverssm`; keeping the old fallback would preserve an unsupported third contract | [v0.29.0 mamba groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704), [fixed-main mamba groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704), [v0.29.0 RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103), [fixed-main RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103) | both, common implementation | Existing constructor UT now asserts the parent-created RecoverSSM value is retained; Ruff and compileall pass | | `worker/v2/model_runner.py`: always forward `kv_cache_allocation_context` | v0.29.0 and fixed main both accept this keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0 signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538), [fixed-main signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566), [vllm-ascend #16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) | both, common implementation | Existing UT continues to assert the exact context object reaches the parent; Ruff and compileall pass | | `_310p/worker/v2/model_runner.py`: select the release KV-zeroing contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes `KVCacheGroupSpec.is_eagle_group` but lacks `SpeculativeConfig.use_eagle_block_drop`; fixed main added the method | vLLM [#53388 diff](https://github.com/vllm-project/vllm/pull/53388/files), commit [`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a); [v0.29.0 group field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200), [fixed-main helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876) | release differs from main | Existing two-path KV-zeroing UT retained and renamed for v0.29.0; Ruff and compileall pass | | existing DFlash kernel UT: remove v0.28-only kwarg omission | the current Ascend kernel accepts the CP arguments and the only excluded lane was v0.28.0, which this PR replaces | [vllm-ascend #15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) | both, common invocation | Existing NPU test remains enabled with all assertions; Current-head CI pending | | Ascend change | Why / upstream cause | Lane | Verification | |---|---|---|---| | `patch/platform/patch_parallel_config.py`, registration, and global patch documentation | Allow Ascend PCP+DP by removing the generic GPU restriction, following vLLM [#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic `parallel_config.current_platform` lookup and rebuild ParallelConfig → SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing configuration/EPLB UTs and historical PCP+DP NPU execution passed; current-head CI pending. | | `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT | Both supported contracts require `is_drafting`, from [#54282](https://github.com/vllm-project/vllm/pull/54282/files), `fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited optimized kernel and use a common wrapper. | Both | Existing positive drafting assertion retained; current-head CI pending. | | `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache UT | Both support `cache_hit_alignment_tokens`, introduced by [#53598](https://github.com/vllm-project/vllm/pull/53598/files), `2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only write-mask branch. | Both | Existing assertions retained; current-head CI pending. | #### Rebase and retired-fallback evidence | Change | Exact source evidence | Decision | |---|---|---| | `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM [#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95), `d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL before v0.29.0; both exact supported sources lack it. The old conditional came from vllm-ascend [#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0 selector and use the existing exclusion for both supported lanes. This does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy skip. | | `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and `tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM [#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29), `12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export `nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The v0.28.0 fallback originated in vllm-ascend [#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete import-failure/`None` fallbacks and the existing UT's obsolete availability skip; assertions remain unchanged and execute on both lanes. | | `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend [#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3), `799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream revert while resolving the real rebase conflict; do not reintroduce the reverted graph-update behavior. | | `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend [#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2), `d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged 310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation branch. | </details> ### Does this PR introduce _any_ user-facing change? Yes. The supported release changes from vLLM v0.28.0 to v0.29.0, and Ascend PCP+DP is enabled on v0.29.0. The fixed main commit remains supported. The vllm-ascend package version is unchanged. Obsolete v0.28-only skips are removed where the contracts are now supported. No skip is migrated to v0.29.0, and no golden data or precision threshold is changed. Existing unrelated exclusions remain exclusions, not passes. ### How was this patch tested? Current-head validation: [run 35420046256](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256) completed successfully on `fd876c5269fb6b39afa04e5a52433505e8786e22`: 39 successful jobs and 6 skipped jobs, no failures. - Pre-commit and mypy passed. Fixed-main [CPU UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105836507531): **5078 passed, 67 skipped, 17 warnings**. Actual vLLM checkout `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, installed `0.1.dev1+g84030bbe3.empty`; Ascend checkout is the PR head. CPU `base_sha` is empty; the log says HEAD is up to date, so no unprinted integration SHA is inferred. - All **32 NPU jobs passed**, 16 per lane. Raw-log session totals per lane: **562 passed, 35 skipped, 1 xfailed** (execution counts, not deduplicated cases). Actual checkouts: main `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, release `98dff2a81d747d1dba01a47f939f48c3526d4206`; installed versions `0.1.dev1+g84030bbe3.empty` and `0.29.0+empty`. Each log confirms Ascend head `fd876c5269fb6b39afa04e5a52433505e8786e22`, integration base `e139b7d573d3769fd1407d5027d7d4831f5469f0`, and HEAD is up to date. - PCP+DP and PCP+PP+MTP suite: [main A3 four-card part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105845372541) and [release A3 four-card part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105845372595) both passed. This does not establish a root cause for historical HCCL failures. - Skipped and expected-failure cases are not passes. **Current-head release CPU verification remains missing**: the existing CPU job runs fixed main only. Historical release CPU results do not fill this gap. No workflow changes or additional skips were introduced. Latest rebase: head `fd876c5269fb6b39afa04e5a52433505e8786e22` onto upstream main `e139b7d573d3769fd1407d5027d7d4831f5469f0`. The old-head green results below are historical; new-head CI has completed successfully; see the verified results below. The fixed vLLM main and release pins remain unchanged. PR-wide Ruff, formatting and Python AST checks passed for 95 Python files; no workflow changes. Rebase reconciliation: - Preserve upstream [#16853](https://github.com/vllm-project/vllm-ascend/pull/16853/files), `8f2e3fed73328148ea603ccfe5764542fa7cbbd2`, without migrating its 0.28-only validator/dispatch gates to 0.29. Both supported pins already call the PCP manager for dispatch token counts. Restore the version-helper import needed by the inherited method. Its inherited 0.28-only skipped tests are not counted as passes. - Adapt the existing unsupported-feature UT to preserve the upstream list: both supported pins delegate PCP checks to the manager via vLLM [#53853](https://github.com/vllm-project/vllm/pull/53853/files), `e376d45e82cb7e220da430e3179e81eb0922cf56`; no obsolete 0.28 string filtering is reinstated. - Adapt the existing KVPP allocation-entry UT from [#15514](https://github.com/vllm-project/vllm-ascend/pull/15514/files), `b64959a11ea69a51409f64e4e68da8e0924dec42`, to the common allocation entry already used by both supported pins after vLLM #51718. Preserve cache-view assertions and the upstream dtype changes. No test functions or skips added. - #16544 remains reverted by #16905; its V4.1 compatibility workaround remains removed. Historical validation before this rebase: - **Current head `8c5d785931c495701bf1da8b5bc80b321ebd3cd0`:** [run 35350119603](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603) completed successfully: **40 successful jobs, 6 skipped jobs, no failures**. Pre-commit and mypy passed. Local Ruff lint/format and AST checks passed for all 95 changed Python files. - **Actual main CPU:** [job 105617611031](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105617611031) reports **5074 passed, 37 skipped**, 17 warnings. Logs confirm the exact PR head, vLLM checkout `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, and installed `0.1.dev1+g84030bbe3.empty`. CPU `base_sha` is empty; no integration merge is inferred. - **Actual dual-version NPU:** All 32 NPU job logs were audited, 16 per lane. Fixed main installed `0.1.dev1+g84030bbe3.empty`; official release checkout `98dff2a81d747d1dba01a47f939f48c3526d4206` installed `0.29.0+empty`. Each lane reports **562 passed, 35 skipped, 1 xfailed** across pytest session summaries (execution counts, not deduplicated unique tests). Logs confirm checkout of this PR head and integration base `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`; the rebase step reports HEAD is up to date. - **PCP combinations:** The context-parallel suite containing PCP+DP and PCP+PP+MTP passes on both [main A3 part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105658864889) and [release A3 part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105658864659). - **Release CPU gap:** Current-head actual release CPU remains missing because the workflow has only a main CPU entry. Version-parameterized UTs and historical release CPU runs do not replace actual release installation. Historical run 35213834751 at head `2560da9c7c2628f1b92b9305a1e80c5856bce74f` passed 4923 tests with 37 skips per lane; its temporary CPU workflow was reverted. - **Backup:** #16898 remains unchanged with no test CI triggered. Local branch `codex/backup-16393-before-16905` preserves previous green `7ad4c82ee`; its results are not used as current-head validation. Skipped/xfail cases and jobs are not counted as passes. Earlier intermittent HCCL errors have no proven root cause; no speculative fix is included. This run passes existing CI, but complete dual-version CPU validation remains outstanding. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
…n V2 runner Extend the layerwise KV pool (AscendStoreConnector + use_layerwise, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to Kimi-K3 hybrid KDA+MLA models on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the @maybe_transfer_kv_layer decorator. Without per-layer hooks the pool worker's current_layer counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips 'thread: 0 save failed' in KVCacheStoreLayerSendingThread. Add the same hooks as ops/gdn.py to the KDA eager-break _forward body: - wait_for_kv_layer_from_connector + record_attention_compute_start after the attn_metadata None check (before conv/recurrent kernels touch mamba state, ordering the deferred per-layer mamba state copy and the layer load) - maybe_save_kv_layer_to_connector on the idle early-exit path and the normal exit path The mamba-hybrid deferral, connector duck-typing and UTs needed by this scenario were merged upstream in vllm-project#15479; this PR only adds the KDA-side hooks. A separate mla_v1 get_kv_cache_shape cache_dtype_str fix that was carried here earlier has been solved upstream by vllm-project#15514. Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer debug checkpoint, TP=8, EP, memcache backend (device_sdma), use_layerwise=true, VLLM_USE_V2_MODEL_RUNNER=1: 104/104 stress requests OK, external prefix cache hit rate 64.6% (90% repeat workload). AIS-Bench A/B vs vLLM HBM prefix caching (input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%) at equal token-level hit rate (85.93%) - the gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover. Signed-off-by: tyy0829 <1455207791@qq.com>
…n V2 runner Extend the layerwise KV pool (AscendStoreConnector + use_layerwise, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to Kimi-K3 hybrid KDA+MLA models on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the @maybe_transfer_kv_layer decorator. Without per-layer hooks the pool worker's current_layer counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips 'thread: 0 save failed' in KVCacheStoreLayerSendingThread. Add the same hooks as ops/gdn.py to the KDA eager-break _forward body: - wait_for_kv_layer_from_connector + record_attention_compute_start after the attn_metadata None check (before conv/recurrent kernels touch mamba state, ordering the deferred per-layer mamba state copy and the layer load) - maybe_save_kv_layer_to_connector on the idle early-exit path and the normal exit path The mamba-hybrid deferral, connector duck-typing and UTs needed by this scenario were merged upstream in vllm-project#15479; this PR only adds the KDA-side hooks. A separate mla_v1 get_kv_cache_shape cache_dtype_str fix that was carried here earlier has been solved upstream by vllm-project#15514. Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer debug checkpoint, TP=8, EP, memcache backend (device_sdma), use_layerwise=true, VLLM_USE_V2_MODEL_RUNNER=1: 104/104 stress requests OK, external prefix cache hit rate 64.6% (90% repeat workload). AIS-Bench A/B vs vLLM HBM prefix caching (input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%) at equal token-level hit rate (85.93%) - the gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover. Signed-off-by: tyy0829 <1455207791@qq.com>
…LA on the V2 model runner (#16320) ### What this PR does / why we need it? Extends the layerwise KV pool (`AscendStoreConnector` + `use_layerwise`, enabled for Qwen3.5 GDN hybrids in #15479) to **Kimi-K3 hybrid KDA + MLA models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the `@maybe_transfer_kv_layer` decorator. Without per-layer hooks the pool worker's `current_layer` counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips `thread: 0 save failed` in `KVCacheStoreLayerSendingThread`. This PR adds the same hooks that `ops/gdn.py` already has to the KDA eager-break `_forward` body (`ops/kimi_kda.py`, +17 lines): - `wait_for_kv_layer_from_connector` + `record_attention_compute_start` right after the `attn_metadata` None check, before the conv/recurrent kernels touch mamba state (this also orders the deferred per-layer mamba state copy and the layer load) - `maybe_save_kv_layer_to_connector` on the idle early-exit path and on the normal exit path **Scope note**: the mamba-hybrid per-layer copy deferral, the connector V1/V2 duck-typing and the accompanying UTs required by this scenario were **merged upstream in #15479** and are not duplicated here. An `mla_v1.get_kv_cache_shape` `cache_dtype_str` signature fix that was briefly carried in this PR was solved upstream by #15514 and has been dropped. ### Does this PR introduce _any_ user-facing change? No API/config change. The layerwise KV pool now works for Kimi-K3 (KDA+MLA hybrid) under the V2 model runner; previously it crashed on the first multi-block prefill. ### How was this patch tested? Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer checkpoint (reduced 4-layer debug build of the ModelSlim W4A8 quantized model), TP=8, EP, eager, memcache backend (`device_sdma`), `use_layerwise=true`, `VLLM_USE_V2_MODEL_RUNNER=1`: - Service starts; `/health` 200; text requests 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (`kvpool hit tokens: 768`); repeat latency 3.4s -> 0.57s - Stress: 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - AIS-Bench A/B vs vLLM HBM prefix caching (prefix dataset: input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): **TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%)** at equal token-level hit rate (85.93% both). The gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover (it only caches the MLA layer KV, the KDA layers still recompute). Remaining: MTP path needs validation once a Kimi-K3 checkpoint with `num_nextn_predict_layers > 0` MTP weights is available (the 4-layer debug checkpoint ships none). - vLLM main: vllm-project/vllm@84030bb Signed-off-by: tyy0829 <1455207791@qq.com>
…bisect Experimental revert to isolate DeepSeek-V4-Pro-w4a8-prefix-cache-PD Nightly performance regression. Do not merge before CI comparison. Signed-off-by: pgzddxx <1697817735@qq.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and vllm-project#15514 --kv-cache-dtype/--indexer_kv_dtype fp8 without the retired enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
…oject#16393) ### What this PR does / why we need it? Upgrade the supported vLLM release dependency from **v0.28.0 to the official v0.29.0**, alongside the fixed vLLM main commit inherited from vllm-project#16216 (already merged). | Reference | Commit | |---|---| | vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` | | PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` | | Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` | | Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a` | | Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` | - Use common implementations where v0.29.0 and the fixed main share KV-cache layouts, Mamba copy/group APIs, PCP handling and speculative-decoding contracts. - Retain explicit `vllm_version_is("0.29.0")` branches for contracts that still differ, including RoPE, scheduler block snapshots, InputBatch, ReplaySSM, KV zeroing and DSpark PP handling. - Remove obsolete v0.28.0 compatibility and adapt existing test fixtures. Version detection uses package versions and the explicit `VLLM_VERSION` override, with local-version suffix handling and UT environment isolation retained; no hard-coded release-SHA inference remains. - Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving other validations and EPLB platform binding. Rebuild dependent Pydantic schemas so nested configuration validation uses the patched validator. The global patch documentation records its rationale and removal criteria. - Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase; use the common exact-release selector and the real function-local DSpark sharing import. This release-only upgrade does not require a new main old-to-new interface scan. The latest rebase incorporates [vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files), which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The now-unnecessary V4.1 drafter import gate and tuple annotation have been removed. Other release adaptations, including vllm-project#15747 Spec+PP handling, remain. Detailed contract evidence is retained below for review. <details> <summary>Per-file adaptations and exact upstream evidence</summary> #### Per-file adaptation ledger Evidence IDs refer to the exact source contract and upstream diff table below. Every row is syntax checked; branch-normalized AST comparison confirms unchanged main function bodies except the version identity helper and the explicitly retired propose argument/type annotations. Both supported versions completed CPU and NPU execution as recorded below. | Ascend file / symbols | Disposition and upstream evidence | Verification | |---|---|---| | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` / `_load_dspark_model_with_target_quant`; `tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased [Ascend vllm-project#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a) (`82df9d871`), including manual PP partition masking. vLLM [#52809 diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share` inside the loader on both supported pins, so retain the earlier release fix: patch `eagle_utils`, never read/patch a nonexistent `dspark_utils._should_share`. v0.29 keeps its global PP guard; main has native PP via [#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks the real function-local import, absent module alias, and restoration on success/failure for both lanes; partition and PP assertions retained. Ruff/syntax pass; actual CPU/NPU pending. | | `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`; `tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29 routing and an obsolete 0.28 dev-build recognition path. For the two supported pins, #50514 exists only on fixed main. Route through `vllm_version_is("0.29.0")`; retain package/local-suffix and explicit environment-override semantics. No release-SHA inference or third release lane. | Existing UTs exercise the real uncached version helper with monkeypatch isolation, local suffix, fixed-main dev string, explicit override, and non-target versions. Isolated routing checks and Ruff/syntax pass; full CPU UT pending. | | `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`, `_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`, `_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436; use shared standardized layouts and retain the v0.29.0 InputBatch gate. The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while removing only legacy v0.28.0 allocation paths. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` | #52839; both lanes use the common PCP import. Rebase keeps current upstream graph code and removes only the v0.28.0 import branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/recompute_scheduler.py`<br>`module imports/dispatch`, `schedule` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`, `AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker` | #52615; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__` | #53614; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module imports/dispatch`, `_get_max_layers_per_page_size`, `_ascend_max_memory_usage_bytes_from_groups`, `_ascend_get_kv_cache_config_from_groups` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` | v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it after KV binding. Evidence: [#52506, `adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c). Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. | Failure reproduced on v0.29.0; Historical release/main NPU validation passed; current-head CI pending | | `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module imports/dispatch` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant` | #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports `_should_share` locally from Eagle utilities; fixed main removes the PP guard. Keep the release `get_pp_group` patch, share through the common Eagle utility, and delete the obsolete v0.28.0 `dspark_utils._should_share` patch. | Exact release failure reproduced; Historical release/main CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`, `__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`, `register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`, `_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`, `_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers | #51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and import the NaN helpers directly because v0.29.0 and fixed main expose the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static checked; historical dual-version CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes `pcp_manager` common to both lanes. The rebase preserves vllm-ascend vllm-project#16409's host-parameter-update revert and removes only the obsolete v0.28.0 capture branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`, `_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/block_table.py`<br>`__init__`, `init_block_table_layout_tensors`, `compute_slot_mappings` | #51718, #51031; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`, `prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`, `execute_model` | #50514, #54436, #52506, #55212, #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch` | #54282; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` | #49811; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module imports/dispatch`, `propose` | #53694, #52188; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`, `wake_up` | #51718, #53508; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | #### Exact upstream evidence | Upstream change | Full commit SHA / direct diff | Actual supported contracts and branch decision | |---|---|---| | #51718 | `8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Both use layers/layer_stride/block_stride/offset, tokens_per_state, CircularBufferSpec and standardized backing; retire shared_by/compress_ratio allocation branches. | | #52839 | `58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8) | Both import PCP operations from vllm.v1.attention.ops.pcp. | | #53896 | `e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec) | Both use Mamba copy-function dictionaries and unwrap UniformTypeKVCacheSpecs. | | #53106 | `1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21) | Both use WeightsMapper instead of AutoWeightsLoader skip_prefixes/skip_substrs. | | #53906 | `98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Only pinned main has the optional MLA storage_block_size dataclass field; release keeps the Ascend derived property, using tokens_per_state. | | #56078 | `719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417) | Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned main uses mrope_num_dims and unified RoPE. | | #52615 | `138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045) | Release uses num_blocks/kv_bytes_per_block; main uses num_chunks/kv_bytes_per_chunk. | | #51358 | `6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release now has boundary_state_offloads and KVConnectorBlockState; remove partial_tail_offloads plumbing. | | #54853 | `0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release constructor takes block_ids snapshots; main takes req_ids and resolve_block_ids. Keep exact release snapshot membership and main lazy-resolution membership. | | #53614 | `144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725) | Only main configures drop_eagle_checkpoint_block for replay-aligned Mamba checkpoints. | | #50514 | `d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Release retains module-level PP/share symbols and Ascend PP workaround; main has the subsequent PP integration. | | #52809 | `91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark module binding to a function-local import from Eagle utilities. The shared Eagle patch remains effective; the old DSpark-module read/write must be removed. | | #54436 | `6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902) | Release InputBatch requires max_seq_len_np; main removed it. | | #52506 | `adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Only main accepts valid_dummy_state_slots/valid_state_slots capture arguments. | | #55212 | `83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Release prepares DCP local sequence lengths before partitioning; main initializes DCP metadata afterwards. | | #53515 | `b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127) | Both accept padded_num_tokens for persistent PCP input buffers. | | #53869 | `b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212) | Both accept pcp_manager during graph capture. | | #51031 | `0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1) | Both distinguish KV and kernel block sizes during DCP slot mapping. | | #54282 | `fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8) | Both gumbel sampling APIs include is_drafting. | | #52188 | `d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0) | Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. | | #53694 | `5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6) | Both propose APIs take DPSyncState; remove the obsolete token-count argument and retain replicated-PCP synchronization. | | #49811 | `01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e) | Both support extract_hidden_states on MRV2; remove old unsupported dispatch/skip. | | #53508 | `479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29) | Both remove post_kv_cache_wake_up; retire release-only call. | | #52494 | `3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1) | Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain release exclusion. | | #52861 | `b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28) | Both include DeepseekV32MTPModel in the two-hidden-state architecture set. | | #54713 | `b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3) | Only main takes replay_boundaries in compressed-prefix hit lookup; preserve release calls without that keyword. | | #42785 | `442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18) | Only main capture callers pass axis_keys; preserve the existing Ascend rejection of nonempty axes. | | #52358 | `8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Both ExecuteModelState have dp_sync; only main has cudagraph_stats. | | #52789 | `9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4) | Release already has mamba_has_prefill_checkpoint_blocks; later main also has fine-grained prefix-cache state. | | #51251 | `7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | v0.29.0 and main expose ec_manager_config; retire the old release-only ScoreEncoder configuration skip. | | #53240 | `b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c) | v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache groups; retire the old release-only replay skip. | | #53853 | `e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both delegate PCP compatibility validation to the PCP manager. | | #53183 | `4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both expose the V1 unsupported-feature helper used by the existing Ascend MRV1 feature filter. | #### Additional inherited contracts | vllm-ascend change | Why it is required | Upstream cause and direct link | Lane | Verification | |---|---|---|---|---| | `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec, list[int]]` from `get_mamba_groups` and both initialize `recoverssm`; keeping the old fallback would preserve an unsupported third contract | [v0.29.0 mamba groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704), [fixed-main mamba groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704), [v0.29.0 RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103), [fixed-main RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103) | both, common implementation | Existing constructor UT now asserts the parent-created RecoverSSM value is retained; Ruff and compileall pass | | `worker/v2/model_runner.py`: always forward `kv_cache_allocation_context` | v0.29.0 and fixed main both accept this keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0 signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538), [fixed-main signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566), [vllm-ascend vllm-project#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) | both, common implementation | Existing UT continues to assert the exact context object reaches the parent; Ruff and compileall pass | | `_310p/worker/v2/model_runner.py`: select the release KV-zeroing contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes `KVCacheGroupSpec.is_eagle_group` but lacks `SpeculativeConfig.use_eagle_block_drop`; fixed main added the method | vLLM [#53388 diff](https://github.com/vllm-project/vllm/pull/53388/files), commit [`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a); [v0.29.0 group field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200), [fixed-main helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876) | release differs from main | Existing two-path KV-zeroing UT retained and renamed for v0.29.0; Ruff and compileall pass | | existing DFlash kernel UT: remove v0.28-only kwarg omission | the current Ascend kernel accepts the CP arguments and the only excluded lane was v0.28.0, which this PR replaces | [vllm-ascend vllm-project#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) | both, common invocation | Existing NPU test remains enabled with all assertions; Current-head CI pending | | Ascend change | Why / upstream cause | Lane | Verification | |---|---|---|---| | `patch/platform/patch_parallel_config.py`, registration, and global patch documentation | Allow Ascend PCP+DP by removing the generic GPU restriction, following vLLM [#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic `parallel_config.current_platform` lookup and rebuild ParallelConfig → SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing configuration/EPLB UTs and historical PCP+DP NPU execution passed; current-head CI pending. | | `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT | Both supported contracts require `is_drafting`, from [#54282](https://github.com/vllm-project/vllm/pull/54282/files), `fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited optimized kernel and use a common wrapper. | Both | Existing positive drafting assertion retained; current-head CI pending. | | `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache UT | Both support `cache_hit_alignment_tokens`, introduced by [#53598](https://github.com/vllm-project/vllm/pull/53598/files), `2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only write-mask branch. | Both | Existing assertions retained; current-head CI pending. | #### Rebase and retired-fallback evidence | Change | Exact source evidence | Decision | |---|---|---| | `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM [#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95), `d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL before v0.29.0; both exact supported sources lack it. The old conditional came from vllm-ascend [vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0 selector and use the existing exclusion for both supported lanes. This does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy skip. | | `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and `tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM [#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29), `12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export `nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The v0.28.0 fallback originated in vllm-ascend [vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete import-failure/`None` fallbacks and the existing UT's obsolete availability skip; assertions remain unchanged and execute on both lanes. | | `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend [vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3), `799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream revert while resolving the real rebase conflict; do not reintroduce the reverted graph-update behavior. | | `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend [vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2), `d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged 310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation branch. | </details> ### Does this PR introduce _any_ user-facing change? Yes. The supported release changes from vLLM v0.28.0 to v0.29.0, and Ascend PCP+DP is enabled on v0.29.0. The fixed main commit remains supported. The vllm-ascend package version is unchanged. Obsolete v0.28-only skips are removed where the contracts are now supported. No skip is migrated to v0.29.0, and no golden data or precision threshold is changed. Existing unrelated exclusions remain exclusions, not passes. ### How was this patch tested? Current-head validation: [run 35420046256](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256) completed successfully on `fd876c5269fb6b39afa04e5a52433505e8786e22`: 39 successful jobs and 6 skipped jobs, no failures. - Pre-commit and mypy passed. Fixed-main [CPU UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105836507531): **5078 passed, 67 skipped, 17 warnings**. Actual vLLM checkout `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, installed `0.1.dev1+g84030bbe3.empty`; Ascend checkout is the PR head. CPU `base_sha` is empty; the log says HEAD is up to date, so no unprinted integration SHA is inferred. - All **32 NPU jobs passed**, 16 per lane. Raw-log session totals per lane: **562 passed, 35 skipped, 1 xfailed** (execution counts, not deduplicated cases). Actual checkouts: main `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, release `98dff2a81d747d1dba01a47f939f48c3526d4206`; installed versions `0.1.dev1+g84030bbe3.empty` and `0.29.0+empty`. Each log confirms Ascend head `fd876c5269fb6b39afa04e5a52433505e8786e22`, integration base `e139b7d573d3769fd1407d5027d7d4831f5469f0`, and HEAD is up to date. - PCP+DP and PCP+PP+MTP suite: [main A3 four-card part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105845372541) and [release A3 four-card part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105845372595) both passed. This does not establish a root cause for historical HCCL failures. - Skipped and expected-failure cases are not passes. **Current-head release CPU verification remains missing**: the existing CPU job runs fixed main only. Historical release CPU results do not fill this gap. No workflow changes or additional skips were introduced. Latest rebase: head `fd876c5269fb6b39afa04e5a52433505e8786e22` onto upstream main `e139b7d573d3769fd1407d5027d7d4831f5469f0`. The old-head green results below are historical; new-head CI has completed successfully; see the verified results below. The fixed vLLM main and release pins remain unchanged. PR-wide Ruff, formatting and Python AST checks passed for 95 Python files; no workflow changes. Rebase reconciliation: - Preserve upstream [vllm-project#16853](https://github.com/vllm-project/vllm-ascend/pull/16853/files), `8f2e3fed73328148ea603ccfe5764542fa7cbbd2`, without migrating its 0.28-only validator/dispatch gates to 0.29. Both supported pins already call the PCP manager for dispatch token counts. Restore the version-helper import needed by the inherited method. Its inherited 0.28-only skipped tests are not counted as passes. - Adapt the existing unsupported-feature UT to preserve the upstream list: both supported pins delegate PCP checks to the manager via vLLM [#53853](https://github.com/vllm-project/vllm/pull/53853/files), `e376d45e82cb7e220da430e3179e81eb0922cf56`; no obsolete 0.28 string filtering is reinstated. - Adapt the existing KVPP allocation-entry UT from [vllm-project#15514](https://github.com/vllm-project/vllm-ascend/pull/15514/files), `b64959a11ea69a51409f64e4e68da8e0924dec42`, to the common allocation entry already used by both supported pins after vLLM #51718. Preserve cache-view assertions and the upstream dtype changes. No test functions or skips added. - vllm-project#16544 remains reverted by vllm-project#16905; its V4.1 compatibility workaround remains removed. Historical validation before this rebase: - **Current head `8c5d785931c495701bf1da8b5bc80b321ebd3cd0`:** [run 35350119603](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603) completed successfully: **40 successful jobs, 6 skipped jobs, no failures**. Pre-commit and mypy passed. Local Ruff lint/format and AST checks passed for all 95 changed Python files. - **Actual main CPU:** [job 105617611031](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105617611031) reports **5074 passed, 37 skipped**, 17 warnings. Logs confirm the exact PR head, vLLM checkout `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, and installed `0.1.dev1+g84030bbe3.empty`. CPU `base_sha` is empty; no integration merge is inferred. - **Actual dual-version NPU:** All 32 NPU job logs were audited, 16 per lane. Fixed main installed `0.1.dev1+g84030bbe3.empty`; official release checkout `98dff2a81d747d1dba01a47f939f48c3526d4206` installed `0.29.0+empty`. Each lane reports **562 passed, 35 skipped, 1 xfailed** across pytest session summaries (execution counts, not deduplicated unique tests). Logs confirm checkout of this PR head and integration base `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`; the rebase step reports HEAD is up to date. - **PCP combinations:** The context-parallel suite containing PCP+DP and PCP+PP+MTP passes on both [main A3 part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105658864889) and [release A3 part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105658864659). - **Release CPU gap:** Current-head actual release CPU remains missing because the workflow has only a main CPU entry. Version-parameterized UTs and historical release CPU runs do not replace actual release installation. Historical run 35213834751 at head `2560da9c7c2628f1b92b9305a1e80c5856bce74f` passed 4923 tests with 37 skips per lane; its temporary CPU workflow was reverted. - **Backup:** vllm-project#16898 remains unchanged with no test CI triggered. Local branch `codex/backup-16393-before-16905` preserves previous green `7ad4c82ee`; its results are not used as current-head validation. Skipped/xfail cases and jobs are not counted as passes. Earlier intermittent HCCL errors have no proven root cause; no speculative fix is included. This run passes existing CI, but complete dual-version CPU validation remains outstanding. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
…LA on the V2 model runner (vllm-project#16320) ### What this PR does / why we need it? Extends the layerwise KV pool (`AscendStoreConnector` + `use_layerwise`, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to **Kimi-K3 hybrid KDA + MLA models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the `@maybe_transfer_kv_layer` decorator. Without per-layer hooks the pool worker's `current_layer` counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips `thread: 0 save failed` in `KVCacheStoreLayerSendingThread`. This PR adds the same hooks that `ops/gdn.py` already has to the KDA eager-break `_forward` body (`ops/kimi_kda.py`, +17 lines): - `wait_for_kv_layer_from_connector` + `record_attention_compute_start` right after the `attn_metadata` None check, before the conv/recurrent kernels touch mamba state (this also orders the deferred per-layer mamba state copy and the layer load) - `maybe_save_kv_layer_to_connector` on the idle early-exit path and on the normal exit path **Scope note**: the mamba-hybrid per-layer copy deferral, the connector V1/V2 duck-typing and the accompanying UTs required by this scenario were **merged upstream in vllm-project#15479** and are not duplicated here. An `mla_v1.get_kv_cache_shape` `cache_dtype_str` signature fix that was briefly carried in this PR was solved upstream by vllm-project#15514 and has been dropped. ### Does this PR introduce _any_ user-facing change? No API/config change. The layerwise KV pool now works for Kimi-K3 (KDA+MLA hybrid) under the V2 model runner; previously it crashed on the first multi-block prefill. ### How was this patch tested? Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer checkpoint (reduced 4-layer debug build of the ModelSlim W4A8 quantized model), TP=8, EP, eager, memcache backend (`device_sdma`), `use_layerwise=true`, `VLLM_USE_V2_MODEL_RUNNER=1`: - Service starts; `/health` 200; text requests 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (`kvpool hit tokens: 768`); repeat latency 3.4s -> 0.57s - Stress: 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - AIS-Bench A/B vs vLLM HBM prefix caching (prefix dataset: input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): **TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%)** at equal token-level hit rate (85.93% both). The gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover (it only caches the MLA layer KV, the KDA layers still recompute). Remaining: MTP path needs validation once a Kimi-K3 checkpoint with `num_nextn_predict_layers > 0` MTP weights is available (the 4-layer debug checkpoint ships none). - vLLM main: vllm-project/vllm@84030bb Signed-off-by: tyy0829 <1455207791@qq.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
…LA on the V2 model runner (vllm-project#16320) ### What this PR does / why we need it? Extends the layerwise KV pool (`AscendStoreConnector` + `use_layerwise`, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to **Kimi-K3 hybrid KDA + MLA models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the `@maybe_transfer_kv_layer` decorator. Without per-layer hooks the pool worker's `current_layer` counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips `thread: 0 save failed` in `KVCacheStoreLayerSendingThread`. This PR adds the same hooks that `ops/gdn.py` already has to the KDA eager-break `_forward` body (`ops/kimi_kda.py`, +17 lines): - `wait_for_kv_layer_from_connector` + `record_attention_compute_start` right after the `attn_metadata` None check, before the conv/recurrent kernels touch mamba state (this also orders the deferred per-layer mamba state copy and the layer load) - `maybe_save_kv_layer_to_connector` on the idle early-exit path and on the normal exit path **Scope note**: the mamba-hybrid per-layer copy deferral, the connector V1/V2 duck-typing and the accompanying UTs required by this scenario were **merged upstream in vllm-project#15479** and are not duplicated here. An `mla_v1.get_kv_cache_shape` `cache_dtype_str` signature fix that was briefly carried in this PR was solved upstream by vllm-project#15514 and has been dropped. ### Does this PR introduce _any_ user-facing change? No API/config change. The layerwise KV pool now works for Kimi-K3 (KDA+MLA hybrid) under the V2 model runner; previously it crashed on the first multi-block prefill. ### How was this patch tested? Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer checkpoint (reduced 4-layer debug build of the ModelSlim W4A8 quantized model), TP=8, EP, eager, memcache backend (`device_sdma`), `use_layerwise=true`, `VLLM_USE_V2_MODEL_RUNNER=1`: - Service starts; `/health` 200; text requests 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (`kvpool hit tokens: 768`); repeat latency 3.4s -> 0.57s - Stress: 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - AIS-Bench A/B vs vLLM HBM prefix caching (prefix dataset: input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): **TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%)** at equal token-level hit rate (85.93% both). The gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover (it only caches the MLA layer KV, the KDA layers still recompute). Remaining: MTP path needs validation once a Kimi-K3 checkpoint with `num_nextn_predict_layers > 0` MTP weights is available (the 4-layer debug checkpoint ships none). - vLLM main: vllm-project/vllm@84030bb Signed-off-by: tyy0829 <1455207791@qq.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
…LA on the V2 model runner (vllm-project#16320) ### What this PR does / why we need it? Extends the layerwise KV pool (`AscendStoreConnector` + `use_layerwise`, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to **Kimi-K3 hybrid KDA + MLA models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the `@maybe_transfer_kv_layer` decorator. Without per-layer hooks the pool worker's `current_layer` counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips `thread: 0 save failed` in `KVCacheStoreLayerSendingThread`. This PR adds the same hooks that `ops/gdn.py` already has to the KDA eager-break `_forward` body (`ops/kimi_kda.py`, +17 lines): - `wait_for_kv_layer_from_connector` + `record_attention_compute_start` right after the `attn_metadata` None check, before the conv/recurrent kernels touch mamba state (this also orders the deferred per-layer mamba state copy and the layer load) - `maybe_save_kv_layer_to_connector` on the idle early-exit path and on the normal exit path **Scope note**: the mamba-hybrid per-layer copy deferral, the connector V1/V2 duck-typing and the accompanying UTs required by this scenario were **merged upstream in vllm-project#15479** and are not duplicated here. An `mla_v1.get_kv_cache_shape` `cache_dtype_str` signature fix that was briefly carried in this PR was solved upstream by vllm-project#15514 and has been dropped. ### Does this PR introduce _any_ user-facing change? No API/config change. The layerwise KV pool now works for Kimi-K3 (KDA+MLA hybrid) under the V2 model runner; previously it crashed on the first multi-block prefill. ### How was this patch tested? Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer checkpoint (reduced 4-layer debug build of the ModelSlim W4A8 quantized model), TP=8, EP, eager, memcache backend (`device_sdma`), `use_layerwise=true`, `VLLM_USE_V2_MODEL_RUNNER=1`: - Service starts; `/health` 200; text requests 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (`kvpool hit tokens: 768`); repeat latency 3.4s -> 0.57s - Stress: 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - AIS-Bench A/B vs vLLM HBM prefix caching (prefix dataset: input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): **TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%)** at equal token-level hit rate (85.93% both). The gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover (it only caches the MLA layer KV, the KDA layers still recompute). Remaining: MTP path needs validation once a Kimi-K3 checkpoint with `num_nextn_predict_layers > 0` MTP weights is available (the 4-layer debug checkpoint ships none). - vLLM main: vllm-project/vllm@84030bb Signed-off-by: tyy0829 <1455207791@qq.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
What this PR does / why we need it?
Part of #13318
This PR refactors the flow so the KV cache quantization dtype is explicitly controlled by the user-facing
--kv-cache-dtype/--attention-config.indexer_kv_dtypearguments at service startup:cache_config.cache_dtype/attention_config.indexer_kv_dtypeviakv_cache_dtype_str_to_dtype()(plus checks for
"fp8"/"int8"), instead of the hardware-profile-based inference indsa_attn_kv_plan/DeepseekV4Indexer, which are removed. - Static quantization becomes mandatory. For GQA/MLA-related models,wrong or missing quantization config now raises a clear
ValueErroron startup instead of silently falling back to unquantized execution (attention_v1.py,quantization/utils.py::enable_fa_quant).enable_sparse_*_c8gating is tied to the actual quantized dtype, so the C8 sparse enter/exit layouts only activate when the dtype is fp8/int8.releases/v0.27.1):worker/patch_kv_cache_dtype.py: flipsSTR_DTYPE_TO_TORCH_DTYPE["fp8"]totorch.float8_e4m3fnin every worker whenregister_kv_cache_dtypeis unavailable (no-op onkvquant_27+).platform/patch_indexer_kv_dtype.py: widens the pydantic LiteralIndexerKVDTypewith"int8"so DeepSeek V4's int8 indexer cache is accepted at config validation time.Does this PR introduce any user-facing change?
Yes — user-facing behavior around quantization becomes explicit:
--kv-cache-dtypeand the indexer cache dtype by--attention-config.indexer_kv_dtype(valuesfp8/int8, previously inferred from the hardware profile).indexer_kv_dtypenow also acceptsint8(Ascend DeepSeek V4 indexer Kcache), which upstream vLLM currently rejects.--kv-cache-dtype fp8|int8and the corresponding quantized weights, vLLM refuses to start with a clear error instead of silently running unquantized.How was this patch tested?